Command reference
Generated from plainml COMMAND --help (run python docs/generate_commands.py to refresh).
plainml
Usage: plainml [OPTIONS] [COMMAND] [ARGS]...
Machine learning in plain English: from a spreadsheet to a trained, explained model.
Quick start:
plainml train data.csv --target price train and compare models
plainml predict latest new.csv -o out.csv predict with the best model
plainml guided mode: answers a few questions
Run 'plainml COMMAND --help' for a command's options and examples.
Options:
-V, --version Show the version and exit.
--debug Show full tracebacks on errors.
-q, --quiet Only print errors.
--no-color Plain text output.
-h, --help Show this message and exit.
Start here:
train Train and compare models; save the best with a report.
profile Summarise a dataset and flag problems, without training.
clean Fix common data problems and save a cleaned copy.
Use a trained model:
predict Predict on new data with a trained model.
evaluate Score a trained model on new labelled data (spot drift).
explain Explain what a model relies on, in plain English.
drift Check whether new data has drifted from the training data.
report Rebuild (and open) a run's HTML report.
Improve a model:
tune Search hyperparameters to squeeze out more accuracy.
importance Rank columns with up to 19 methods; find the ones that matter.
select Keep only the columns that matter; save a smaller dataset.
Other kinds of problem:
cluster Group similar rows (no target needed).
anomaly Find unusual rows (fraud, errors, outliers).
forecast Forecast a value over time (sales, demand, traffic...).
Deploy and share:
web Open the plainml website: upload data, run anything, download results.
serve Serve a model as a REST API (FastAPI).
deploy Package a model as a Docker image that serves it.
export Export a model to ONNX or MLflow format.
Housekeeping:
runs List past runs, or delete old ones.
compare Compare runs side by side.
models List the models plainml can train.
init Create a starter config file.
Start here
plainml train
Usage: plainml train [OPTIONS] [DATA]
Train many models, rank them with cross-validation, and save the best one.
DATA is a CSV/Excel/Parquet/JSON file, a URL, or a database URL.
Options:
-t, --target TEXT Column to predict. Repeat or comma-separate for several.
--task [auto|classification|regression]
Override the auto-detected task.
-m, --metric TEXT Metric to rank models by (default: f1 or rmse). E.g. accuracy,
roc_auc, r2, mae.
--models TEXT Only try these models, e.g. rf,xgboost. See: plainml models
--exclude TEXT Skip these models.
--quick Only fast models, no ensemble.
--thorough Also try the extra models (AdaBoost, Gaussian process, PyTorch...)
and a stacked ensemble.
--cv INTEGER RANGE Cross-validation folds. [default: 5] [2<=x<=20]
--test-size FLOAT RANGE Share of rows held out for the final test. [default: 0.2]
[0.05<=x<=0.5]
--seed INTEGER Random seed for reproducible results. [default: 42]
--time-budget DURATION Stop starting new models after this long, e.g. 90s, 5m, 1h.
--balance [auto|none|weights|smote]
Handle imbalanced classes. [default: auto]
--threshold [auto|on|off] Tune the yes/no decision threshold. [default: auto]
--log-target Model log(target): helps with skewed, positive amounts like
prices.
--calibrate Adjust predicted probabilities so '70% sure' really means right
70% of the time.
--mlflow Also log the run to MLflow (the server in MLFLOW_TRACKING_URI, or
a local store).
--ensemble / --no-ensemble Also try averaging the top 3 models. [default: on]
--refit / --no-refit Retrain the winner on all rows before saving. [default: on]
--save-all Also save every model, not just the best.
--zip Also zip the run folder.
--drop COLUMNS Columns to ignore.
--keep COLUMNS Columns to use even if they look like IDs.
--sample N Train on a random sample: a row count, or a fraction like 0.1.
--n-jobs INTEGER CPU cores to use (-1 = all). [default: -1]
-o, --out DIR Where to save runs. [default: runs]
--name TEXT Name for the run folder.
--report / --no-report Write report.html. [default: on]
--private Keep raw data out of the report, model and run folder (for sharing
results).
--open Open the report when done.
-c, --config FILE YAML file of options (see: plainml init).
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
Examples:
plainml train churn.csv --target churned
plainml train houses.xlsx -t price --metric mae --quick
plainml train data.csv -t label --models rf,lightgbm,xgboost --time-budget 10m
plainml train data.csv -t label --drop customer_id --balance smote
plainml train reviews.csv -t is_spam,is_urgent (multi-label)
plainml train --config plainml.yaml
plainml profile
Usage: plainml profile [OPTIONS] DATA
Show each column's type, gaps and values, plus warnings (imbalance, leaks, IDs...).
Examples:
plainml profile sales.csv
plainml profile churn.csv --target churned --html profile.html --open
Options:
-t, --target TEXT Also check this target column (imbalance, leakage, task).
--drop COLUMNS Columns to ignore.
--keep COLUMNS Columns to use even if they look like IDs.
--html FILE Also write an HTML version.
--open Open the HTML version.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml clean
Usage: plainml clean [OPTIONS] DATA
Standardise blanks, fix numbers stored as text, parse dates, remove duplicates and more.
Examples:
plainml clean raw.csv -o clean.csv
plainml clean raw.xlsx --impute --outliers clip --target price
Options:
-o, --output FILE Where to save the cleaned data. [default: DATA_clean.csv]
-t, --target TEXT Target column: rows missing it are dropped; it's never imputed or
encoded.
--duplicates / --keep-duplicates
Remove duplicate rows. [default: duplicates]
--impute / --no-impute Fill blanks (median / most common value). [default: no-impute]
--outliers [none|clip|remove] Handle extreme numeric values (1.5×IQR rule). [default: none]
--encode One-hot encode text categories (for tools that need numbers).
--drop-ids / --keep-ids Remove ID-like and constant columns. [default: drop-ids]
--dry-run Show what would change without saving.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
Use a trained model
plainml predict
Usage: plainml predict [OPTIONS] MODEL [DATA]
Predict with MODEL: a .joblib file, a run folder, part of a run name, or 'latest'.
Examples:
plainml predict latest new_customers.csv -o predictions.csv
plainml predict runs/20260924-101500_churn new.xlsx --proba
plainml predict latest --horizon 30 (forecast models)
plainml predict latest huge.csv -o out.csv --chunk-size 200000
Options:
-o, --output FILE Save predictions (.csv, .xlsx, .json, .parquet).
--proba Add a probability column per class.
--strict Fail if any input column is missing (instead of treating it as
blank).
--horizon INTEGER Forecast models: how many periods ahead.
--chunk-size INTEGER RANGE Stream a big CSV/Parquet file this many rows at a time (needs -o).
[x>=1000]
--show INTEGER Rows to print. [default: 10]
--runs-dir TEXT Where runs are saved. [default: runs]
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml evaluate
Usage: plainml evaluate [OPTIONS] MODEL DATA
Check how MODEL does on DATA (which must include the target column).
Compares against the scores measured at training time, so you can tell when the world has changed
and the model needs retraining.
Example:
plainml evaluate latest march_labelled.csv --report march.html
Options:
--report FILE Also write an HTML report.
--runs-dir TEXT Where runs are saved. [default: runs]
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml explain
Usage: plainml explain [OPTIONS] [MODEL] [DATA]
Which columns matter, which way they push predictions, and why one row got its prediction.
Uses the run's held-out test rows unless you pass DATA.
Examples:
plainml explain
plainml explain latest new.csv --row 3
plainml explain runs/20260924-101500_churn --shap
Options:
--top INTEGER How many columns to show. [default: 15]
--shap Also compute SHAP values (pip install "plainml[explain]").
--row INTEGER Explain the prediction for this row number (0-based) of DATA.
-o, --output FILE Save the importance table.
--runs-dir TEXT Where runs are saved. [default: runs]
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml drift
Usage: plainml drift [OPTIONS] REFERENCE DATA
Compare DATA with a model's training data (REFERENCE = 'latest', a run, a .joblib) or with an
older data file. No labels needed.
Examples:
plainml drift latest this_month.csv
plainml drift january.csv june.csv --target churned
Options:
-t, --target TEXT Target column to leave out (when REFERENCE is a data file).
--out, --runs-dir TEXT Where runs are found (for REFERENCE) and this one is saved.
[default: runs]
--name TEXT Name for the run folder.
--report / --no-report Write report.html. [default: report]
--open Open the report when done.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml report
Usage: plainml report [OPTIONS] [RUN]
Regenerate report.html for RUN ('latest', a folder, or part of a run name).
Options:
--open / --no-open [default: open]
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
Improve a model
plainml tune
Usage: plainml tune [OPTIONS] [SOURCE]
Tune the best models from a run (default: latest), or models on a data file.
Uses Optuna when installed (pip install "plainml[tune]"), otherwise random search.
Examples:
plainml tune
plainml tune latest --trials 100 --timeout 20m
plainml tune data.csv -t price --models lightgbm,rf
Options:
-t, --target TEXT Target column (when SOURCE is a data file).
--models TEXT Models to tune (default: the run's best models).
--top INTEGER Tune this many of the run's best models. [default: 3]
--trials INTEGER Settings to try per model. [default: 30]
--timeout DURATION Overall time limit, e.g. 10m.
-m, --metric TEXT Metric to optimise (default: the run's).
--cv INTEGER RANGE Cross-validation folds. [2<=x<=20]
--seed INTEGER Random seed.
--open Open the report when done.
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
plainml importance
Usage: plainml importance [OPTIONS] DATA
Which columns matter for predicting the target, measured many different ways.
Filters (mutual information, F-test, correlation...), model-based importances (random forest,
L1...), model-agnostic ones (permutation, drop-column, SHAP) and searches (RFE, forward/backward
selection, exhaustive search, Boruta, stability selection) are combined into one consensus
ranking, with a recommended set and a report.
Examples:
plainml importance data.csv -t price
plainml importance data.csv -t churned --methods all --open
plainml importance data.csv -t churned --methods boruta,exhaustive -o reduced.csv
Options:
-t, --target TEXT Column to predict. [required]
--methods TEXT Methods (comma-separated), or 'all'. Default: mutual_info, f_test,
correlation, random_forest, l1, permutation, rfe, boruta. Also:
chi2, variance, extra_trees, gradient_boosting, linear,
drop_column, shap, forward, backward, exhaustive, stability.
-k, --keep-top INTEGER How many columns to recommend (default: those that clearly help).
-o, --output FILE Also save DATA with only the recommended columns (+ target).
--drop COLUMNS Columns to ignore.
--sample N Use a random sample: a row count, or a fraction like 0.1.
--seed INTEGER Random seed. [default: 42]
--out TEXT Where to save the run. [default: runs]
--name TEXT Name for the run folder.
--private Keep raw values out of the report.
--open Open the report when done.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml select
Usage: plainml select [OPTIONS] DATA
Like 'importance', without saving a run: rank the columns and keep the best ones.
Examples:
plainml select data.csv -t price
plainml select data.csv -t churned -k 10 -o reduced.csv
Options:
-t, --target TEXT Column to predict. [required]
-k, --keep-top INTEGER How many columns to keep (default: those that clearly help).
--methods TEXT Methods (comma-separated), or 'all'. Default: mutual_info, f_test,
correlation, random_forest, l1, permutation, rfe, boruta. Also:
chi2, variance, extra_trees, gradient_boosting, linear,
drop_column, shap, forward, backward, exhaustive, stability.
-o, --output FILE Save DATA with only the selected columns (+ target).
--drop COLUMNS Columns to ignore.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
Other kinds of problem
plainml cluster
Usage: plainml cluster [OPTIONS] DATA
Find natural groups (segments) in DATA and describe what makes each one different.
Examples:
plainml cluster customers.csv
plainml cluster customers.csv -k 4 --drop customer_id -o segments.csv
Options:
-k, --clusters TEXT Number of groups, a range like 2-8, or auto. [default: auto]
--algorithms TEXT Subset of: kmeans, agglomerative, gmm, hdbscan, kmedoids.
--drop COLUMNS Columns to ignore.
-o, --output FILE Save DATA with a 'cluster' column.
--seed INTEGER Random seed. [default: 42]
--out TEXT Where to save the run. [default: runs]
--name TEXT Name for the run folder.
--private Keep raw values out of the report and run folder.
--open Open the report when done.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml anomaly
Usage: plainml anomaly [OPTIONS] DATA
Score every row by how unusual it is and explain what's odd about the top ones.
Examples:
plainml anomaly transactions.csv -o scored.csv
plainml anomaly transactions.csv --label is_fraud --contamination 0.01
Options:
--contamination TEXT Expected share of anomalies, e.g. 0.02, or auto. [default: auto]
--label TEXT Optional column marking known anomalies (1/0), used to score the
detectors.
--algorithms TEXT Subset of: iforest, lof, ocsvm, robust.
--drop COLUMNS Columns to ignore.
-o, --output FILE Save DATA with anomaly scores.
--seed INTEGER Random seed. [default: 42]
--out TEXT Where to save the run. [default: runs]
--name TEXT Name for the run folder.
--private Keep raw values out of the report and run folder.
--open Open the report when done.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
plainml forecast
Usage: plainml forecast [OPTIONS] DATA
Backtest several forecasting models on DATA's history and forecast the future.
Rows dated after the last known target value are read as plans for --inputs, so a promo calendar
can sit in the same file as the history.
Examples:
plainml forecast sales.csv -t revenue --horizon 30
plainml forecast visits.csv -t visitors --date day --freq D -o next_month.csv
plainml forecast stores.csv -t sales --group store --inputs promo --country US
plainml forecast sales.csv -t sales --inputs promo,price --future plans.csv
Options:
-t, --target TEXT Column to forecast. [required]
--date TEXT Date/time column (auto-detected if left out).
--horizon INTEGER How many periods ahead (default: depends on frequency).
--freq TEXT Frequency: D (daily), W, M, H... (auto-detected if left out).
--agg [sum|mean|last] How to combine several rows with the same date. [default: sum]
--models TEXT Subset of: naive, seasonal_naive, ridge, rf, histgb, lightgbm,
xgboost.
--group TEXT Forecast each value of this column separately (store, product,
region...).
--inputs TEXT Columns known in advance that affect the target (promo, price...).
Comma-separated.
--future FILE File with the inputs' planned future values (date, inputs, and
group if used).
--country TEXT Add public holidays for this country code (US, GB, IN, DE...).
-o, --output FILE Save the forecast.
--seed INTEGER Random seed. [default: 42]
--out TEXT Where to save the run. [default: runs]
--name TEXT Name for the run folder.
--open Open the report when done.
--sheet NAME Excel sheet to read (default: the largest, named in a warning).
--query SQL SQL query to run (with a database URL).
--table NAME Database table to read (with a database URL).
--engine [auto|pandas|polars] Reader for big CSV/Parquet files. [default: auto]
-h, --help Show this message and exit.
Deploy and share
plainml web
Usage: plainml web [OPTIONS]
Start a local website for everything plainml does.
Drag in a file, choose a task (predict a column, forecast, find groups or anomalies, rank columns,
check drift, profile or clean), watch it run, then read the report and download any result file.
Examples:
plainml web
plainml web --port 9000 --runs-dir projects/churn/runs
plainml web --host 0.0.0.0 --token change-me (share on your network)
plainml web --export site (static, runs in the browser)
Options:
--host TEXT Use 0.0.0.0 to let other machines connect (add --token). [default:
127.0.0.1]
--port INTEGER Port to listen on. [default: 8765]
--runs-dir TEXT Where runs are saved. [default: runs]
--token TEXT Require this access token to use the site (or set
PLAINML_WEB_TOKEN).
--max-upload-mb INTEGER RANGE Largest upload. [default: 500; x>=1]
--no-browser Don't open a browser tab.
--export DIRECTORY Instead of starting a server, write a static version that runs
plainml in the visitor's browser (for Vercel, GitHub Pages...).
-h, --help Show this message and exit.
plainml serve
Usage: plainml serve [OPTIONS] [MODEL]
Start a prediction API for MODEL, with interactive docs at /docs.
Example:
plainml serve latest --port 8000
curl -X POST localhost:8000/predict -H 'Content-Type: application/json' \
-d '{"rows": [{"age": 42, "plan": "pro"}]}'
Options:
--host TEXT Use 0.0.0.0 to accept outside connections. [default: 127.0.0.1]
--port INTEGER Port to listen on. [default: 8000]
--api-key TEXT Require this key in the X-API-Key header (or set PLAINML_API_KEY).
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
plainml deploy
Usage: plainml deploy [OPTIONS] [MODEL]
Write a folder with a Dockerfile, pinned requirements and MODEL, ready to build.
Example:
plainml deploy latest
docker build -t churn deploy/20260101-120000_churn
docker run -p 8000:8000 -e PLAINML_API_KEY=secret churn
Options:
-o, --output DIRECTORY Folder to write. [default: deploy/<run>]
--port INTEGER Port to listen on. [default: 8000]
--python TEXT Python version for the image, e.g. 3.12.
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
plainml export
Usage: plainml export [OPTIONS] [MODEL]
Convert MODEL to ONNX (checked against the original) or save it as an MLflow model.
Examples:
plainml export latest -o churn.onnx
plainml export latest --format mlflow -o churn_mlflow
Options:
--format [onnx|mlflow] onnx: runs outside Python; mlflow: a folder for MLflow's model registry.
[default: onnx]
-o, --output FILE Output file. [default: next to the model]
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
Housekeeping
plainml runs
Usage: plainml runs [OPTIONS]
List saved runs, newest last.
Examples:
plainml runs
plainml runs --prune --keep 10
plainml runs --prune --older-than 30d --kind drift
Options:
--runs-dir TEXT Where runs are saved. [default: runs]
-n, --limit INTEGER Runs to list. [default: 20]
--prune Delete old runs (asks first). Use with --keep / --older-than.
--keep INTEGER RANGE With --prune: keep the newest N runs. [x>=0]
--older-than AGE With --prune: only runs older than this, e.g. 30d, 2w.
--kind TEXT With --prune: only this kind of run (train, forecast, importance...).
-y, --yes Don't ask for confirmation.
-h, --help Show this message and exit.
plainml compare
Usage: plainml compare [OPTIONS] [RUN_REFS]...
Compare RUNs (default: the last five): data, target, best model and scores.
Examples:
plainml compare
plainml compare runs/*churn*
Options:
--runs-dir TEXT Where runs are saved. [default: runs]
-h, --help Show this message and exit.
plainml models
Usage: plainml models [OPTIONS]
Show every model, which tasks it supports, and whether it's installed.
Options:
--task [classification|regression]
Only models for this task.
-h, --help Show this message and exit.
plainml init
Usage: plainml init [OPTIONS] [PATH]
Write a commented YAML config you can edit and run with: plainml train --config PATH
Options:
--data TEXT Data file to put in the config.
-t, --target TEXT Target column to put in the config.
-h, --help Show this message and exit.