The concept
Data pipelines you can see, approve and explain.
PipeCraft is built on one conviction: the value of data work is not speed, it is trust. Everything in the platform exists to make a transformation visible, deliberate and traceable.
Philosophy
A pipeline nobody can read is a pipeline nobody can trust.
Data teams rarely fail because a transformation is hard. They fail because the transformation is invisible: a script on someone’s laptop, a silent filter, a fix applied once and forgotten.
PipeCraft makes each of those things explicit. Rules are nodes on a canvas. Invalid rows go to a named place. A change needs a new approval. A run leaves a report. Nothing happens by accident.
Visible
One lane per column, one node per rule. If you can read a flowchart, you can read a PipeCraft pipeline.
Governed
A human approves the configuration before it can run. The approver and the time are written into the flow.
Traceable
Outputs, quarantine and a run report are produced every time, so every row has a story.
The engine
Eight stages, always in the same order
Predictable order is what makes a pipeline reviewable. Pick a stage to see what it does.
Load
load_source_step
Reads the source into a table: a file (CSV, Parquet, JSON, NDJSON), a REST API or a database. Files are scanned for viruses on upload, and database or HTTP sources pass an SSRF guard.
- InputA source definition
- OutputThe raw dataset
Mapping
apply_mapping_step
Renames columns and aligns the dataset with the target schema — types, required columns and column order — before any cleaning begins.
schema_: columns: age: { type: int, required: true }
Rules
apply_rules_step
The heart of the pipeline. Each column carries its own chain of rules: cleaning (trim, fill nulls, regex), casting, date parsing, numeric transforms, filters and derived columns. The catalogue holds over a hundred.
columns: email: rules: [ trim, lower, { not_null }, { regex: "^[^@]+@[^@]+$" } ]
Scalers
apply_scalers_step
Scales numeric columns when the data is destined for machine learning. Five scalers are available: standard, minmax, robust, maxabs and log1p.
LLM
llm_column_step
An optional, strictly bounded stage. The model sees a deterministic sample and may only answer with a supported action (no-op, regex replace, value map). Anything else is ignored and traced.
Quarantine
quarantine_step
Splits valid rows from invalid ones. Invalid rows are kept, with their reason, in a separate file. If their share exceeds the configured threshold, the run stops before writing anything.
- Default threshold1 %
- Setting
stop_if_ratio_gt
Outputs
write_outputs_step
Writes the clean dataset and the quarantine file to files (CSV, Parquet, JSON) and/or to a database table, which you can then explore with the built-in SQL browser.
Finalize
finalize_step
Closes the run and assembles the run report: what was read, what each stage changed, what was quarantined and where the outputs went.
Governance
Nothing runs without a human “yes”
A pipeline lives in one of two states. Validation produces a draft; approval turns it into something that may run.
- Explicit approval. The approver’s identity and the timestamp are stored in the flow.
- Change means re-approve. Edit the configuration and you start a new cycle. No silent drift.
- Real data, real profile. The draft is generated by profiling the actual source, not a guess.
Quarantine
Bad rows are evidence, not garbage
Try it: decide how many rows fail and how tolerant the pipeline is.
Medallion layers
Raw, refined, ready
For document workflows, data is organised in three layers so that you can always go back to the source.
Bronze
The untouched input, stored as received. Your safety net and your audit trail.
Silver
Cleaned, typed and validated data — the layer analysts and pipelines build on.
Gold
Business-ready datasets, shaped for dashboards and models.
Crafty
An AI co-pilot with guard-rails
Describe the outcome — “clean postal codes”, “cast age to int and drop nulls” — and Crafty edits the canvas for you from the Ctrl+K command bar.
Crafty is constrained by the same rule catalogue as everyone else, and it changes a draft. A human still reviews and approves the result.
Trim names, lowercase emails and drop rows without an email.
Done — 3 rules added on 2 columns. Review the draft and approve when ready.
After the pipeline
Trusted data is only useful if it goes somewhere
BI Studio
Drag-and-drop dashboards on the clean output, with optional public share links.
ML Studio
Train scikit-learn classification or regression models and keep metrics and artifacts per pipeline.
Database Studio
Query the tables your pipeline wrote, straight from the browser.
Who it is for
One platform, three ways to use it
Analysts
Clean and reshape data without waiting for engineering.
Data engineers
YAML flows, a remote CLI and CI-friendly runs.
Organisations
Roles, shared pipelines and an audit trail across the whole team.
FAQ
Questions we hear often
Do I need to know how to code?
No. The canvas, the rule palette and Crafty cover day-to-day cleaning. Engineers can still work with the YAML file when they prefer.
What happens to rows that fail a rule?
They are written to a quarantine file with the reason, not deleted. If too many fail, the run stops so you can investigate.
Can the AI change my pipeline on its own?
No. Crafty edits a draft using catalogue rules only. A person must approve the configuration before it can run.
Which sources and formats are supported?
CSV, Parquet, JSON and NDJSON files up to 20 GB, REST APIs and databases.
Can I run pipelines from CI?
Yes. The remote CLI logs in with a device code, then lets you list, pull, push and run pipelines and follow their status.
See your own data go from raw to trusted
Every new account starts with a ready-made demo workspace: a pipeline, a model and a dashboard.
Create an account