A multi-page Streamlit application for exploring CSV datasets, comparing baseline machine-learning models, exporting trained models, and running document Q&A, emotion analysis, zero-shot classification, and grammar correction.
Features · Workflow · AutoML · NLP · Installation · Troubleshooting
AutoMLapp brings common early-stage data-science workflows into one browser interface. A user can upload a CSV file, inspect its schema and data quality, create interactive visualizations, compare baseline models with PyCaret, download the selected model, and experiment with several NLP pipelines.
The application is designed for:
- rapid dataset assessment;
- educational AutoML demonstrations;
- baseline modelling before deeper experimentation;
- lightweight NLP prototyping;
- portfolio demonstrations of end-to-end Streamlit workflows.
No coding is required to use the interface after installation.
- Upload any standard CSV dataset
- Inspect row count, column count, missing values, and duplicates
- Preview the first 100 records
- Review data types, non-null counts, missing percentages, and cardinality
- Generate descriptive statistics across numeric and categorical columns
- Explore numeric distributions with histograms and box plots
- Visualize numeric correlations with an interactive heatmap
- Review the 20 most frequent categorical values
- Inspect missing values by column
- Build Plotly scatter plots without writing code
- Select numeric X and Y axes
- Colour observations by any available column
- Zoom, pan, hover, and export through Plotly controls
- Compare regression models with PyCaret
- Compare classification models with PyCaret
- Run DBSCAN clustering experiments
- Display the selected estimator and complete score table
- Inspect estimator parameters
- Use a fixed session seed (
42) for repeatable setup behavior - Save and download the resulting
.pklmodel
- Retrieval-augmented question answering over PDF documents
- Emotion/sentiment classification with EmoRoBERTa
- Zero-shot text classification with BART-MNLI
- Grammar correction with a T5 model
flowchart TD
A["Upload CSV"] --> B["Local dataset file"]
B --> C["EDA dashboard"]
B --> D["Interactive Plotly charts"]
B --> E["PyCaret experiments"]
E --> F["Scores + model export"]
G["PDF or text"] --> H["NLP utilities"]
| Page | Purpose | Output |
|---|---|---|
| Upload dataset | Read and store a CSV file | Data preview and local dataset.csv |
| Data analysis | Profile structure, statistics, and quality | Metrics, tables, distributions, and correlations |
| Data visualisation | Explore numeric relationships | Interactive Plotly scatter plot |
| ML models | Run supervised or unsupervised baselines | Selected model, leaderboard, scores, and parameters |
| Download model | Export the trained estimator | best_model.pkl |
| NLP | Run one of four language tasks | Answer, label, emotion, or corrected text |
The built-in dashboard replaces the heavier YData Profiling integration used by earlier versions of the project. This avoids known pandas compatibility problems while keeping the most useful exploratory views directly inside Streamlit.
- first 100 rows;
- column types;
- non-null and missing counts;
- missing percentage;
- unique-value count;
- descriptive statistics.
- selectable numeric feature;
- 30-bin histogram;
- outlier-aware box plot;
- full correlation matrix when at least two numeric columns exist.
- selectable categorical feature;
- missing values represented explicitly;
- horizontal bar chart of the 20 most frequent values.
- missing-value counts by column;
- duplicate-row count;
- clear success messages when no issue is detected.
Choose Regression when the target is continuous and numeric. PyCaret prepares the data, evaluates its supported regression estimators, and selects the best baseline under its comparison configuration.
Choose Classification when the target represents discrete classes, including text labels such as Setosa, Versicolor, or Fraud.
Choose Clustering to fit a DBSCAN model without a supervised target. The application displays the clustering score table returned by PyCaret and exports the fitted pipeline.
Every PyCaret setup call uses:
session_id=42The chosen estimator is saved through PyCaret as:
best_model.pkl
Important
AutoML comparison produces a baseline, not a production-ready model. Validate the target, preprocessing, leakage risk, cross-validation design, business metric, fairness, and out-of-sample performance before deployment.
After modelling completes, the interface displays:
- selected estimator class;
- readable estimator representation;
- PyCaret comparison or clustering metrics;
- complete estimator parameter table;
- local model path.
The Download model page makes the serialized pipeline available as best_model.pkl.
Caution
Python pickle files can execute code while loading. Only load model artifacts that you created or obtained from a trusted source, and recreate the original dependency environment whenever possible.
The document assistant builds a temporary retrieval pipeline for an uploaded PDF:
flowchart TD
A["Uploaded PDF"] --> B["PyPDF page loader"]
B --> C["1,024-character chunks"]
C --> D["MPNet embeddings"]
D --> E["scikit-learn vector store"]
F["User question"] --> E
E --> G["Top 3 chunks"]
G --> H["Falcon-7B-Instruct"]
H --> I["Generated answer"]
| Component | Implementation |
|---|---|
| PDF loading | LangChain PyPDFLoader |
| Chunking | 1,024 characters with 64-character overlap |
| Embeddings | sentence-transformers/all-mpnet-base-v2 |
| Retrieval | SKLearnVectorStore, top 3 chunks |
| Hosted language model | tiiuae/falcon-7b-instruct through Hugging Face Hub |
| Chain | RetrievalQA with stuff context combination |
This task requires a Hugging Face Hub token and network access.
The interface uses:
arpanghoshal/EmoRoBERTa
The pipeline returns the highest-scoring emotion label and confidence percentage for the supplied sentence.
The user provides a sentence and comma-separated candidate labels. The application uses:
facebook/bart-large-mnli
It returns the highest-ranked label and its confidence score without task-specific fine-tuning.
The grammar assistant uses Happy Transformer with:
vennify/t5-base-grammar-correction
Generation runs with five-beam search and returns a corrected sentence.
.
├── app.py # Streamlit entry point
├── help.py # Backward-compatible PDF-QA wrapper
├── src/automl_app/
│ ├── __init__.py
│ ├── app.py # Pages, state, modelling, and NLP UI
│ └── services/
│ ├── __init__.py
│ └── document_qa.py # PDF retrieval-QA pipeline
├── .gitignore
├── requirements.txt
└── README.md
The following runtime artifacts are intentionally excluded from Git:
dataset.csv
uploaded_file.pdf
best_model.pkl
document_vector_db.parquet
.env
- Python 3.11 recommended
pip- A modern browser
- Internet access for initial NLP model downloads
- A Hugging Face token only for document Q&A
The PyCaret, Numba, and scientific Python versions are intentionally pinned. Python 3.11 provides the safest installation path for this dependency set.
git clone https://github.com/younesgu/AutoMLapp.git
cd AutoMLapp# Windows PowerShell
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1# macOS or Linux
python3.11 -m venv .venv
source .venv/bin/activatepython -m pip install --upgrade pip
python -m pip install --prefer-binary -r requirements.txtThe first installation may take several minutes because PyCaret includes a broad scientific Python stack.
python -m streamlit run app.pyOpen http://localhost:8501 if Streamlit does not open automatically.
Create .env in the repository root:
HUGGINGFACEHUB_API_TOKEN=your_hugging_face_tokenRestart Streamlit after saving the token.
The embedding model runs through Sentence Transformers, while the final question-answering prompt and retrieved PDF context are sent to the configured Hugging Face hosted model.
Warning
Do not use document Q&A with confidential PDFs unless external processing through Hugging Face has been explicitly approved by your organization.
The other NLP models are downloaded on first use and cached by the Hugging Face libraries. Downloads can require significant time, memory, and disk space.
- Open Upload dataset.
- Upload a
.csvfile. - Open Data analysis to inspect quality and distributions.
- Use Data visualisation for interactive relationships.
- Upload the dataset.
- Open ML models.
- Select the target column.
- Choose Regression for a continuous numeric target or Classification for class labels.
- Click Run modelling.
- Review the leaderboard, selected model, and parameters.
- Open Download model to export the pipeline.
- Upload a dataset containing meaningful modelling features.
- Open ML models and choose Clustering.
- Click Run modelling to fit DBSCAN.
- Review the returned cluster metrics and exported model.
- Open NLP.
- Select the task.
- Provide the required PDF, text, labels, or question.
- Run the task and review the result.
AutoMLapp writes working artifacts to the repository directory:
| File | Contents | Lifecycle |
|---|---|---|
dataset.csv |
Most recently uploaded dataset | Replaced by the next CSV upload |
uploaded_file.pdf |
Most recently uploaded PDF | Replaced by the next PDF upload |
best_model.pkl |
Most recently saved PyCaret model | Replaced by later modelling runs |
These files are ignored by Git but remain on the local filesystem until replaced or deleted. Do not assume that uploaded information disappears when the browser tab closes.
When scientific Python packages conflict, recreating only the project environment is usually the safest fix:
# Windows PowerShell
deactivate
Remove-Item -Recurse -Force .venv
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install --prefer-binary -r requirements.txt| Error or symptom | Likely cause | Solution |
|---|---|---|
Matplotlib requires numpy>=... |
Packages were installed independently with conflicting versions | Recreate .venv and install only from requirements.txt |
NDFrame.infer_objects() got an unexpected keyword argument 'copy' |
Incompatible pandas/YData Profiling versions from an older environment | Recreate .venv; the current app no longer uses YData Profiling |
'DataFrame' object has no attribute 'profile_report' |
Old application code or missing profiling integration | Pull the current version; analysis is now built into render_analysis |
could not convert string to float: 'Setosa' |
Regression was selected for a categorical target | Choose Classification for discrete labels |
streamlit is not recognized |
The virtual environment is inactive | Activate .venv or run python -m streamlit run app.py |
| Hugging Face token error | HUGGINGFACEHUB_API_TOKEN is missing |
Add the token to .env and restart Streamlit |
| NLP task appears frozen | A large model is downloading or loading | Check network activity and wait for the first model load |
| Memory error during modelling | The dataset or model comparison exceeds available RAM | Sample the data, reduce features, or use a machine with more memory |
The repository currently has no automated test suite. A lightweight syntax check can be run with:
python -m compileall app.py help.py srcRecommended next checks include unit tests for CSV handling, dataset requirements, feature selection, model persistence, and mocked NLP services.
- CSV files are loaded fully into memory.
- Uploaded datasets and PDFs are stored in the project directory.
- The application supports one local working dataset and model at a time.
- PyCaret model comparison can be slow on large datasets or CPU-only machines.
- DBSCAN parameters are currently fixed by the PyCaret default workflow.
- The interface does not expose preprocessing, validation, or cross-validation controls.
- NLP models are large and can require substantial RAM and disk space.
- Document Q&A depends on a hosted model and an external token.
- Generated answers and zero-shot labels can be inaccurate.
- There is no authentication, user isolation, persistent experiment tracking, or automated test suite.
- Confirm that uploaded datasets and PDFs may legally be processed.
- Remove sensitive or personally identifiable information when possible.
- Review model outputs for leakage, bias, and class imbalance.
- Never deploy the selected baseline without independent validation.
- Treat NLP output as assistance, not verified fact.
- Keep API tokens in
.envand never commit them to Git.