google/bigquery-bigframes
>- Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.
npx skills add https://github.com/google/skills --skill bigquery-bigframes
BigFrames is a Python library that lets you take advantage of BigQuery
data processing by using familiar Python APIs.
via BigFrames methods to leverage BigQuery's scale rather than downloading
data.
importing BigFrames. This speeds up data processing significantly by relaxing
row-sequence constraints.
import bigframes.pandas as bpd
bpd.options.bigquery.ordering_mode = 'partial'
peek() for data preview: Use peek(n) to preview data instead ofhead(n). peek(n) randomly samples n rows and is significantly faster.
head(n) returns rows in strict order and fails in partial ordering mode
unless the DataFrame has been explicitly sorted.
to_pandas() download alldata to client memory, bypassing BigQuery’s distributed computation and
risking Out of Memory (OOM) errors. Do not materialize data locally unless:
read_gbq() if a DataFrame/Series method achieves the same result, as it
breaks the Pandas abstraction and prevents lazy query execution.
df.col.str.*, df.col.dt.*) instead ofremote User Defined Functions (UDFs). UDFs require extra resources and
time to deploy.
Series.map() or DataFrame.apply(). Thesemethods do not accept functions without udf or remote_function
decorators.
# Avoid:
df["upper"] = df["name"].map(lambda x: x.upper())
# Prefer:
df["upper"] = df["name"].str.upper()
Proactively verify schemas using .dtypes and inspect sample records using
display() with .peek().
possible. BigFrames is compatible with Matplotlib and Seaborn. If direct
plotting fails, use the .plot accessor. If the dataset is too large to plot,
aggregate or sample the data before calling
.to_pandas() to plot locally.
bigframes.bigquery.ml package: Do not use Scikit-learn or other MLlibraries with BigQuery DataFrames. Standard Scikit-learn models require
bringing data into local client memory, whereas bigframes.bigquery.ml
delegates training directly to BigQuery's scalable ML engine. Import functions
from bigframes.bigquery.ml.
regression model to predict numerical values.
regression model to predict boolean values.
The BigFrames ML package (bigframes.ml) is a legacy package that mimics the
scikit-learn API but is no longer recommended for new projects. Only use this
package if the user explicitly requests BigFrames ML.
classes from bigframes.ml instead of bigframes.bigquery.ml.
predict() method always returns a DataFrame containing both predictions
and features, rather than a single series of predictions.
random_state: Do not pass a random_state argument wheninstantiating BigFrames ML models, as this parameter is not supported in the
BigFrames ML package.
OneHotEncoder or StandardScaler unlessexplicitly requested, as scaling is handled automatically.
BigFrames lacks GridSearchCV or RandomizedSearchCV.
bigframes.ml.forecasting.horizon.
transform() method. Use predict()instead.
model.to_gbq(). To load apersisted model, use bpd.read_gbq_model().
Take google/bigquery-bigframes from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.