diff --git a/README.md b/README.md new file mode 100644 index 0000000..4a25ff5 --- /dev/null +++ b/README.md @@ -0,0 +1,250 @@ +# RECLASS: Multi-Task Deep Learning for App Review Classification +**RECLASS** tests whether a single shared encoder with four task-specific heads +can match four independently trained single-task models performance, while producing all +four predictions in one forward pass instead of four. + +> Trained model checkpoints, datasets, full experiment logs, and extended commit history\ +> are mirrored on:\ +> [git.charlierasberry.dev/mik/ReClass](https://git.charlierasberry.dev/mik/ReClass) + +## The four tasks + +From a single review, the model predicts: + +| Task | Type | Labels | +|---|---|---| +| Bug Report | binary | Yes / No | +| Feature Request | binary | Yes / No | +| Aspect | 6-class | App, Driver, General, Payment, Pricing, Service | +| Aspect Sentiment | 3-class | Positive, Neutral, Negative | + +Example: *"Got charged twice for one ride, please fix"* → bug: yes, feature: no, +aspect: payment, sentiment: negative. + +## Architecture + +![Architecture](architecture.png) + +- **Shared encoder:** XLM-RoBERTa-base (12 layers, 768 hidden, 12 attention heads). + Multilingual encoder chosen over BERT-base after data exploration showed + substantial language switching in individual reviews. +- **Tokenisation:** SentencePiece, `max_length=256` (chosen after measuring + truncation impact on minority-class reviews at shorter lengths). +- **Shared regularisation:** dropout (p = 0.2) applied to the shared `[CLS]` + representation before branching into task specific heads. +- **Task heads:** four independent linear layers projecting the 768-dim + representation into each task's class space\ + (Bug: 768→2, Feature: 768→2, + Aspect: 768→6, Sentiment: 768→3). +- **Loss:** unweighted sum of four inverse-frequency-weighted cross-entropy + losses, per task. +- **Optimiser:** AdamW, lr 2e-5, weight decay 0.1, linear warmup (10% of steps), + gradient clipping at max_norm 1.0. +- **Early stopping:** on validation macro-F1, patience of 3 epochs. + +> Single task baselines use the same encoder as its MTL equivalent, the only difference is the number of heads. + +## Headline result: MTL matches STL + +McNemar's exact test on paired predictions (750 test samples per task, α = 0.05): + +| Task | STL Macro-F1 | MTL Macro-F1 | Δ | p-value | Significant? | +|---|---|---|---|---|---| +| Bug Report | 0.785 | 0.784 | −0.001 | 0.699 | No | +| Feature Request | 0.742 | 0.763 | +0.021 | 0.653 | No | +| Aspect | 0.694 | 0.717 | +0.023 | 0.210 | No | +| Aspect Sentiment | 0.786 | 0.758 | −0.028 | 0.728 | No | + +No task shows a statistically significant difference. MTL is competitive across +the board and slightly ahead on two of four tasks, while using only one +encoder (~278M params), a 4× +reduction from (\~1.1B params), and producing all four predictions in a single forward pass instead +of four. + +## The keyword-boosted experiment - tradeoffs and wins + +A second dataset was curated using keyword sampling for minority classes - bug reports and feature requests to test whether it is worthwhile. + +| Task | MTL Original | MTL Boosted | Δ | +|---|---|---|---| +| Bug Report | 0.783 | 0.905 | **+0.122** | +| Feature Request | 0.763 | 0.816 | +0.053 | +| Aspect | 0.717 | 0.803 | +0.086 | +| Aspect Sentiment | 0.757 | 0.600 | **−0.157** | + +Three tasks improved substantially. Aspect Sentiment collapsed. Bug-report +keyword reviews were heavily negative skewed, so oversampling them shifted the training +distribution to 80.9% negative sentiment, in an MTL scenario this negatively affects the ability to learn the shared representations. + +> **In a shared-encoder MTL setup, data curation for one task is never just +> curation for that task alone.** Shared Learning is also shared interference. + +*(The models were not cross-evaluated - see Limitations.)* + +## Dataset + +Manually annotated from the [Ola Vs Uber Reviews](https://www.kaggle.com/datasets/khushipitroda/ola-vs-uber-play-store-reviews) dataset +using a custom Tkinter annotation tool (`multitag.py`) with keyboard shortcuts +and enforced label completeness. + +| Stage | Count | +|---|---| +| Raw corpus | 1,069,616 | +| After cleaning + ≥5-word filter | 495,036 | +| Original distribution, annotated | 4,999 | +| Keyword-boosted, annotated | 4,997 | +| Split | 70% train / 15% val / 15% test | + +Class imbalance is severe in the wild: bug reports are 18.4% of the original +set, neutral sentiment just 5.2%. Handled via inverse-frequency class weighting +during training, not by rebalancing the raw data. + +**Data availability:** the annotated datasets (original and keyword-boosted, +train/val/test splits) are included in this repository under `data/`. Raw +source: [Uber Customer Reviews on Kaggle](https://www.kaggle.com/datasets/rajatraj0502/ola-vs-uber-reviews). + +## Repository structure + +``` +. +├── architecture.png # Overview diagram +├── environment.yml # conda environment spec +├── data/ # annotated original + keyword-boosted splits +├── notebooks/ # preprocessing, annotation QA, analysis, inference prep +├── src/ +│ ├── preprocess.py # cleaning + filtering pipeline +│ ├── sampler.py # original / keyword-boosted sampling strategies +│ ├── multitag.py # Tkinter annotation tool +│ ├── dataset.py # tokenisation + PyTorch Dataset +│ ├── model.py # shared-encoder MTL + single-task model definitions +│ ├── train.py # training loop (MTL + STL) +│ ├── evaluate.py # Macro-F1, confusion matrices, McNemar's test +│ └── infer.py # inference on new review text +└── README.md +``` + +## Running the pipeline + +### 0. Optional Setup and get data +```bash +cd # replace with location +mkdir -p data/ +mkdir -p data/raw/ +mkdir -p data/processed/ +mkdir -p outputs/ +mkdir -p runs/ # if using tensorboard +mkdir -p outputs/inference/ + +# Optionally replace with own data though could cause slight issues +curl -L -o data/raw/ola-vs-uber-play-store-reviews.zip\ + https://www.kaggle.com/api/v1/datasets/download/khushipitroda/ola-vs-uber-play-store-reviews + +cd data/raw +unzip ola-vs-uber-play-store-reviews.zip +mv "Uber Customer Reviews.csv" uber_reviews.csv +rm ola-vs-uber-play-store-reviews.zip +rm “Ola Customer Reviews.csv” +cd ../.. + +conda env create -f environment.yml +conda activate reclass + +``` + + + +### 1. Preprocess raw reviews (clean, filter, sample) +```bash +python src/preprocess.py [optionally fill in INPUT and OUTPUT at top of file] +``` + + +### 2. Sample cleaned reviews for tagging +```bash +python src/sampler.py [optionally fill in target_samples, INPUT and OUTPUT at top of file] +``` + + +### 3. Annotate (or use existing labelled CSVs) +```bash +python src/multitag.py [optionally fill in INPUT and OUTPUT at top of file] +``` + + +### 3.5 Train Test Split +Use the notebook preprocessing_tagged.ipynb to create splits, +cells 2 and 23 contain input/output paths. +- Inputs data/raw/uber_reviews_tagged_original.csv and /data/raw/uber_reviews_tagged_boosted.csv by default. +- Outputs to data/tagged_boosted_cleaned.csv and data/tagged_original_cleaned.csv by default. + + +### 4. Train - MTL or single-task baseline +```bash +python src/train.py +Defaults to python src/train.py --mode mtl --task all --dataset original –batch_size 16 --epochs 5 --patience 3 --lr 2e-5 +``` + + +### 5. Evaluate against test set +```bash +python src/evaluate.py + +options: + -h, --help show this help message and exit + --mode {mtl,stl} mtl or stl + --task {all,bug_report,feature_request,aspect,aspect_sentiment} + --dataset {original,boosted} + --model_path MODEL_PATH .pt file path + --batch_size BATCH_SIZE + +e.g. python src/evaluate.py --mode mtl --dataset original --model_path outputs/best_model_mtl_original.pt +``` + + +### 6. Run inference on new text +```bash +python src/infer.py +options: + -h, --help show this help message and exit + --model MODEL .pt file - just the filename not path + --task {all,bug_report,feature_request,aspect,aspect_sentiment} + --interactive Loops reading input until exit() + --text Use command line text for input + --dataset DATASET Enter a file name for inference (stored in data/processed/) + --batch_size BATCH_SIZE + --mode {mtl,stl} mtl or stl + --text_column TEXT_COLUMN Where is the text column + +e.g. python src/infer.py --model best_model_mtl_original.pt --mode mtl --text +``` + + +>Trained on a single NVIDIA RTX 2070 Super (8GB VRAM), ~12 GPU-hours total across all reported runs. + +## Limitations + +- **Single annotator.** No inter-annotator agreement / Cohen's Kappa, so inconsistencies / bias can't be ruled out. +- **No cross-evaluation.** The boosted model was only tested on its own + (curated - not original distribution) test set, so its performance on the real data was left unknown. +- **Single domain.** Cross domain performance was left unknown. +- **Overconfidence on errors.** Overconfidence was too high and needs calibrating in the future. +- **Linear classification heads.** Tasks with multiple classes would likely benefit from an additional layer. + +## Future work + +- Cross-evaluate the boosted model on the original test set, quickest next step. +- Second annotator on any future data labelling. +- Dynamic or learned weighting. +- Possibly implement the transferrable elements elsewhere on future projects. +- Temperature scaling on the overconfidence. + +## Citation / provenance + +Built as a BSc Computer Science dissertation project, Oxford Brookes University +(2026). Uses the [Ola Vs Uber Reviews](https://www.kaggle.com/datasets/khushipitroda/ola-vs-uber-play-store-reviews) also uploaded here [Ola Vs Uber Play store reviews](https://www.kaggle.com/datasets/rajatraj0502/ola-vs-uber-reviews) +from Kaggle, and [XLM-RoBERTa](https://arxiv.org/abs/1911.02116) +(Conneau et al., 2020) via HuggingFace Transformers. + +## License + +*[MIT - see LICENSE](LICENSE)*