Clean your data
from the terminal.
DataPrepX is an interactive, terminal-based preprocessing studio that turns raw CSV and Excel files into analysis-ready data — with a local LLM that can inspect and edit your dataset on demand.
Everything you need to ship
clean datasets, fast.
A no-code workflow with a power-user soul. Configure once, reproduce forever.
CSV & Excel Loading
Drop in .csv, .xlsx or .xls files via flag or interactive prompt with tab completion.
Manual Cell Editor
Fix specific cells before the pipeline runs. Overwrite, set to NaN, or skip — with context preview.
Reproducible Pipeline
Six configurable stages run in a fixed order with a live progress bar and full audit trail.
Local LLM Assistant
Ollama-powered chat that can query and edit your DataFrame using natural language.
Multi-Format Export
Save processed data to CSV, JSON and XLSX simultaneously — plus a job_config.json snapshot.
Target Protection
Detach your label column before transformations and reattach safely after — never accidentally encoded.
Rich Terminal UX
Coloured panels, bordered tables, spinners and arrow-key menus. Built on rich + questionary.
Sandboxed Execution
LLM pandas expressions are blocklisted from os, sys, eval — your shell stays safe.
Streaming Responses
LLM replies stream token-by-token to the terminal for a responsive, real-time feel.
Six stages. Zero surprises.
Every transformation runs in a fixed, reproducible order. Configure each step once, re-run forever.
Target Protection
Detach the label column before all transforms and reattach safely afterward.
Missing Values
Seven strategies for handling nulls, from statistical imputation to fills and drops.
Categorical Encoding
Convert text columns to numbers using sklearn encoders or pandas get_dummies.
Outlier Handling
Detect and remove anomalous rows using classic Tukey fences or robust covariance.
Feature Selection
Drop near-zero-variance features that carry little predictive signal.
Feature Scaling
Normalise numeric columns so their ranges are comparable across models.
Your dataset,
in conversation.
DataPrepX integrates a local Ollama LLM that can inspect, query and even edit your DataFrame. Available before preprocessing — and again after — so you can explore, fix, and verify your data without ever leaving the terminal.
Query in plain English
"What is the average Age?" — the LLM runs pandas under the hood and replies in natural language.
Edit cells via chat
"Fix the Age in row 12 — it should be 34." Dtype-aware casting handles ints, floats, bools, and NaN.
Full audit log
Every LLM-driven edit is recorded. Type 'exit' to see a clean table of column, row, old → new.
Sandboxed & local
Runs entirely on your machine via Ollama. A blocklist prevents access to os, sys, eval and friends.
Up and running in 30 seconds.
Clone the repository
Run the installer
Run DataPrepX
Open source. Built for builders.
DataPrepX is Apache 2.0 licensed. Fork it, extend it, ship your own pipeline steps — we're actively reviewing PRs.