v1.0 — open source & LLM-powered

Clean your data
from the terminal.

DataPrepX is an interactive, terminal-based preprocessing studio that turns raw CSV and Excel files into analysis-ready data — with a local LLM that can inspect and edit your dataset on demand.

Explore features
dataprepx — cli.py
$ python cli.py --file sales.csv
─────────────────────────────────────────────
┃ DataPrepX — Smart Data Preprocessing Studio
─────────────────────────────────────────────
✓ Loaded sales.csv (12,483 rows × 14 cols)
? Preview the dataset? › Yes
? Pick target column: › revenue
? Missing strategy: › median
? Encode categorical: › one-hot
? Scale features: › standard
⠿ Running pipeline ━━━━━━━━━━━━━━━━━━━━ 100% • 0:02
✓ Saved processed_sales.csv
✓ Saved job_config.json
$
// features

Everything you need to ship
clean datasets, fast.

A no-code workflow with a power-user soul. Configure once, reproduce forever.

CSV & Excel Loading

Drop in .csv, .xlsx or .xls files via flag or interactive prompt with tab completion.

Manual Cell Editor

Fix specific cells before the pipeline runs. Overwrite, set to NaN, or skip — with context preview.

Reproducible Pipeline

Six configurable stages run in a fixed order with a live progress bar and full audit trail.

Local LLM Assistant

Ollama-powered chat that can query and edit your DataFrame using natural language.

Multi-Format Export

Save processed data to CSV, JSON and XLSX simultaneously — plus a job_config.json snapshot.

Target Protection

Detach your label column before transformations and reattach safely after — never accidentally encoded.

Rich Terminal UX

Coloured panels, bordered tables, spinners and arrow-key menus. Built on rich + questionary.

Sandboxed Execution

LLM pandas expressions are blocklisted from os, sys, eval — your shell stays safe.

Streaming Responses

LLM replies stream token-by-token to the terminal for a responsive, real-time feel.

// pipeline

Six stages. Zero surprises.

Every transformation runs in a fixed, reproducible order. Configure each step once, re-run forever.

01
step 01

Target Protection

Detach the label column before all transforms and reattach safely afterward.

any column
02
step 02

Missing Values

Seven strategies for handling nulls, from statistical imputation to fills and drops.

meanmedianmost_frequentconstantffillbfilldrop
03
step 03

Categorical Encoding

Convert text columns to numbers using sklearn encoders or pandas get_dummies.

LabelOne-HotOrdinalNone
04
step 04

Outlier Handling

Detect and remove anomalous rows using classic Tukey fences or robust covariance.

IQREllipticEnvelopeNone
05
step 05

Feature Selection

Drop near-zero-variance features that carry little predictive signal.

0.000.010.050.100.20
06
step 06

Feature Scaling

Normalise numeric columns so their ranges are comparable across models.

StandardMinMaxRobustNone
// llm assistant

Your dataset,
in conversation.

DataPrepX integrates a local Ollama LLM that can inspect, query and even edit your DataFrame. Available before preprocessing — and again after — so you can explore, fix, and verify your data without ever leaving the terminal.

Query in plain English

"What is the average Age?" — the LLM runs pandas under the hood and replies in natural language.

Edit cells via chat

"Fix the Age in row 12 — it should be 34." Dtype-aware casting handles ints, floats, bools, and NaN.

Full audit log

Every LLM-driven edit is recorded. Type 'exit' to see a clean table of column, row, old → new.

Sandboxed & local

Runs entirely on your machine via Ollama. A blocklist prevents access to os, sys, eval and friends.

dataprepx — chat (llama3.2)
you › Which row has the highest Salary?
tool query_data → df['Salary'].idxmax()
llama › Row 847 — employee "M. Chen" earns $184,200.
you › Mark the Score in row 0 as missing.
tool edit_cell → Score[0] = NaN
llama › Done. Score in row 0 is now NaN.
📋 Edit Audit Log
ColumnRowOldNewScore087.3NaN
// quickstart

Up and running in 30 seconds.

1

Clone the repository

bash
$ git clone https://github.com/shreyas23dev/DataPrepX-A-tool-for-Enhancing-quality-through-pre-processing-technique-in-Data-Analytics.git
2

Run the installer

bash
$ sh installer.sh # installs deps & sets up Ollama
3

Run DataPrepX

bash
$ python cli.py # interactive file prompt
$ python cli.py --file my_dataset.csv

Open source. Built for builders.

DataPrepX is Apache 2.0 licensed. Fork it, extend it, ship your own pipeline steps — we're actively reviewing PRs.