Upload files to "/"

This commit is contained in:
mik
2026-07-17 18:59:25 +00:00
parent fc3eea7a26
commit c344983320

250
README.md Normal file
View File

@@ -0,0 +1,250 @@
# RECLASS: Multi-Task Deep Learning for App Review Classification
**RECLASS** tests whether a single shared encoder with four task-specific heads
can match four independently trained single-task models performance, while producing all
four predictions in one forward pass instead of four.
> Trained model checkpoints, datasets, full experiment logs, and extended commit history\
> are mirrored on:\
> [git.charlierasberry.dev/mik/ReClass](https://git.charlierasberry.dev/mik/ReClass)
## The four tasks
From a single review, the model predicts:
| Task | Type | Labels |
|---|---|---|
| Bug Report | binary | Yes / No |
| Feature Request | binary | Yes / No |
| Aspect | 6-class | App, Driver, General, Payment, Pricing, Service |
| Aspect Sentiment | 3-class | Positive, Neutral, Negative |
Example: *"Got charged twice for one ride, please fix"* → bug: yes, feature: no,
aspect: payment, sentiment: negative.
## Architecture
![Architecture](architecture.png)
- **Shared encoder:** XLM-RoBERTa-base (12 layers, 768 hidden, 12 attention heads).
Multilingual encoder chosen over BERT-base after data exploration showed
substantial language switching in individual reviews.
- **Tokenisation:** SentencePiece, `max_length=256` (chosen after measuring
truncation impact on minority-class reviews at shorter lengths).
- **Shared regularisation:** dropout (p = 0.2) applied to the shared `[CLS]`
representation before branching into task specific heads.
- **Task heads:** four independent linear layers projecting the 768-dim
representation into each task's class space\
(Bug: 768→2, Feature: 768→2,
Aspect: 768→6, Sentiment: 768→3).
- **Loss:** unweighted sum of four inverse-frequency-weighted cross-entropy
losses, per task.
- **Optimiser:** AdamW, lr 2e-5, weight decay 0.1, linear warmup (10% of steps),
gradient clipping at max_norm 1.0.
- **Early stopping:** on validation macro-F1, patience of 3 epochs.
> Single task baselines use the same encoder as its MTL equivalent, the only difference is the number of heads.
## Headline result: MTL matches STL
McNemar's exact test on paired predictions (750 test samples per task, α = 0.05):
| Task | STL Macro-F1 | MTL Macro-F1 | Δ | p-value | Significant? |
|---|---|---|---|---|---|
| Bug Report | 0.785 | 0.784 | 0.001 | 0.699 | No |
| Feature Request | 0.742 | 0.763 | +0.021 | 0.653 | No |
| Aspect | 0.694 | 0.717 | +0.023 | 0.210 | No |
| Aspect Sentiment | 0.786 | 0.758 | 0.028 | 0.728 | No |
No task shows a statistically significant difference. MTL is competitive across
the board and slightly ahead on two of four tasks, while using only one
encoder (~278M params), a 4×
reduction from (\~1.1B params), and producing all four predictions in a single forward pass instead
of four.
## The keyword-boosted experiment - tradeoffs and wins
A second dataset was curated using keyword sampling for minority classes - bug reports and feature requests to test whether it is worthwhile.
| Task | MTL Original | MTL Boosted | Δ |
|---|---|---|---|
| Bug Report | 0.783 | 0.905 | **+0.122** |
| Feature Request | 0.763 | 0.816 | +0.053 |
| Aspect | 0.717 | 0.803 | +0.086 |
| Aspect Sentiment | 0.757 | 0.600 | **0.157** |
Three tasks improved substantially. Aspect Sentiment collapsed. Bug-report
keyword reviews were heavily negative skewed, so oversampling them shifted the training
distribution to 80.9% negative sentiment, in an MTL scenario this negatively affects the ability to learn the shared representations.
> **In a shared-encoder MTL setup, data curation for one task is never just
> curation for that task alone.** Shared Learning is also shared interference.
*(The models were not cross-evaluated - see Limitations.)*
## Dataset
Manually annotated from the [Ola Vs Uber Reviews](https://www.kaggle.com/datasets/khushipitroda/ola-vs-uber-play-store-reviews) dataset
using a custom Tkinter annotation tool (`multitag.py`) with keyboard shortcuts
and enforced label completeness.
| Stage | Count |
|---|---|
| Raw corpus | 1,069,616 |
| After cleaning + ≥5-word filter | 495,036 |
| Original distribution, annotated | 4,999 |
| Keyword-boosted, annotated | 4,997 |
| Split | 70% train / 15% val / 15% test |
Class imbalance is severe in the wild: bug reports are 18.4% of the original
set, neutral sentiment just 5.2%. Handled via inverse-frequency class weighting
during training, not by rebalancing the raw data.
**Data availability:** the annotated datasets (original and keyword-boosted,
train/val/test splits) are included in this repository under `data/`. Raw
source: [Uber Customer Reviews on Kaggle](https://www.kaggle.com/datasets/rajatraj0502/ola-vs-uber-reviews).
## Repository structure
```
.
├── architecture.png # Overview diagram
├── environment.yml # conda environment spec
├── data/ # annotated original + keyword-boosted splits
├── notebooks/ # preprocessing, annotation QA, analysis, inference prep
├── src/
│ ├── preprocess.py # cleaning + filtering pipeline
│ ├── sampler.py # original / keyword-boosted sampling strategies
│ ├── multitag.py # Tkinter annotation tool
│ ├── dataset.py # tokenisation + PyTorch Dataset
│ ├── model.py # shared-encoder MTL + single-task model definitions
│ ├── train.py # training loop (MTL + STL)
│ ├── evaluate.py # Macro-F1, confusion matrices, McNemar's test
│ └── infer.py # inference on new review text
└── README.md
```
## Running the pipeline
### 0. Optional Setup and get data
```bash
cd <project location> # replace with location
mkdir -p data/
mkdir -p data/raw/
mkdir -p data/processed/
mkdir -p outputs/
mkdir -p runs/ # if using tensorboard
mkdir -p outputs/inference/
# Optionally replace with own data though could cause slight issues
curl -L -o data/raw/ola-vs-uber-play-store-reviews.zip\<newline>
https://www.kaggle.com/api/v1/datasets/download/khushipitroda/ola-vs-uber-play-store-reviews
cd data/raw
unzip ola-vs-uber-play-store-reviews.zip
mv "Uber Customer Reviews.csv" uber_reviews.csv
rm ola-vs-uber-play-store-reviews.zip
rm “Ola Customer Reviews.csv”
cd ../..
conda env create -f environment.yml
conda activate reclass
```
### 1. Preprocess raw reviews (clean, filter, sample)
```bash
python src/preprocess.py [optionally fill in INPUT and OUTPUT at top of file]
```
### 2. Sample cleaned reviews for tagging
```bash
python src/sampler.py [optionally fill in target_samples, INPUT and OUTPUT at top of file]
```
### 3. Annotate (or use existing labelled CSVs)
```bash
python src/multitag.py [optionally fill in INPUT and OUTPUT at top of file]
```
### 3.5 Train Test Split
Use the notebook preprocessing_tagged.ipynb to create splits,
cells 2 and 23 contain input/output paths.
- Inputs data/raw/uber_reviews_tagged_original.csv and /data/raw/uber_reviews_tagged_boosted.csv by default.
- Outputs to data/tagged_boosted_cleaned.csv and data/tagged_original_cleaned.csv by default.
### 4. Train - MTL or single-task baseline
```bash
python src/train.py
Defaults to python src/train.py --mode mtl --task all --dataset original batch_size 16 --epochs 5 --patience 3 --lr 2e-5
```
### 5. Evaluate against test set
```bash
python src/evaluate.py
options:
-h, --help show this help message and exit
--mode {mtl,stl} mtl or stl
--task {all,bug_report,feature_request,aspect,aspect_sentiment}
--dataset {original,boosted}
--model_path MODEL_PATH .pt file path
--batch_size BATCH_SIZE
e.g. python src/evaluate.py --mode mtl --dataset original --model_path outputs/best_model_mtl_original.pt
```
### 6. Run inference on new text
```bash
python src/infer.py
options:
-h, --help show this help message and exit
--model MODEL .pt file - just the filename not path
--task {all,bug_report,feature_request,aspect,aspect_sentiment}
--interactive Loops reading input until exit()
--text Use command line text for input
--dataset DATASET Enter a file name for inference (stored in data/processed/)
--batch_size BATCH_SIZE
--mode {mtl,stl} mtl or stl
--text_column TEXT_COLUMN Where is the text column
e.g. python src/infer.py --model best_model_mtl_original.pt --mode mtl --text
```
>Trained on a single NVIDIA RTX 2070 Super (8GB VRAM), ~12 GPU-hours total across all reported runs.
## Limitations
- **Single annotator.** No inter-annotator agreement / Cohen's Kappa, so inconsistencies / bias can't be ruled out.
- **No cross-evaluation.** The boosted model was only tested on its own
(curated - not original distribution) test set, so its performance on the real data was left unknown.
- **Single domain.** Cross domain performance was left unknown.
- **Overconfidence on errors.** Overconfidence was too high and needs calibrating in the future.
- **Linear classification heads.** Tasks with multiple classes would likely benefit from an additional layer.
## Future work
- Cross-evaluate the boosted model on the original test set, quickest next step.
- Second annotator on any future data labelling.
- Dynamic or learned weighting.
- Possibly implement the transferrable elements elsewhere on future projects.
- Temperature scaling on the overconfidence.
## Citation / provenance
Built as a BSc Computer Science dissertation project, Oxford Brookes University
(2026). Uses the [Ola Vs Uber Reviews](https://www.kaggle.com/datasets/khushipitroda/ola-vs-uber-play-store-reviews) also uploaded here [Ola Vs Uber Play store reviews](https://www.kaggle.com/datasets/rajatraj0502/ola-vs-uber-reviews)
from Kaggle, and [XLM-RoBERTa](https://arxiv.org/abs/1911.02116)
(Conneau et al., 2020) via HuggingFace Transformers.
## License
*[MIT - see LICENSE](LICENSE)*