RECLASS: Multi-Task Deep Learning for App Review Classification

RECLASS tests whether a single shared encoder with four task-specific heads can match four independently trained single-task models performance, while producing all four predictions in one forward pass instead of four.

Trained model checkpoints, datasets, full experiment logs, and extended commit history
are mirrored on:
git.charlierasberry.dev/mik/ReClass

The four tasks

From a single review, the model predicts:

Task Type Labels
Bug Report binary Yes / No
Feature Request binary Yes / No
Aspect 6-class App, Driver, General, Payment, Pricing, Service
Aspect Sentiment 3-class Positive, Neutral, Negative

Example: "Got charged twice for one ride, please fix" → bug: yes, feature: no, aspect: payment, sentiment: negative.

Architecture

Architecture

  • Shared encoder: XLM-RoBERTa-base (12 layers, 768 hidden, 12 attention heads). Multilingual encoder chosen over BERT-base after data exploration showed substantial language switching in individual reviews.
  • Tokenisation: SentencePiece, max_length=256 (chosen after measuring truncation impact on minority-class reviews at shorter lengths).
  • Shared regularisation: dropout (p = 0.2) applied to the shared [CLS] representation before branching into task specific heads.
  • Task heads: four independent linear layers projecting the 768-dim representation into each task's class space
    (Bug: 768→2, Feature: 768→2, Aspect: 768→6, Sentiment: 768→3).
  • Loss: unweighted sum of four inverse-frequency-weighted cross-entropy losses, per task.
  • Optimiser: AdamW, lr 2e-5, weight decay 0.1, linear warmup (10% of steps), gradient clipping at max_norm 1.0.
  • Early stopping: on validation macro-F1, patience of 3 epochs.

Single task baselines use the same encoder as its MTL equivalent, the only difference is the number of heads.

Headline result: MTL matches STL

McNemar's exact test on paired predictions (750 test samples per task, α = 0.05):

Task STL Macro-F1 MTL Macro-F1 Δ p-value Significant?
Bug Report 0.785 0.784 0.001 0.699 No
Feature Request 0.742 0.763 +0.021 0.653 No
Aspect 0.694 0.717 +0.023 0.210 No
Aspect Sentiment 0.786 0.758 0.028 0.728 No

No task shows a statistically significant difference. MTL is competitive across the board and slightly ahead on two of four tasks, while using only one encoder (~278M params), a 4× reduction from (~1.1B params), and producing all four predictions in a single forward pass instead of four.

The keyword-boosted experiment - tradeoffs and wins

A second dataset was curated using keyword sampling for minority classes - bug reports and feature requests to test whether it is worthwhile.

Task MTL Original MTL Boosted Δ
Bug Report 0.783 0.905 +0.122
Feature Request 0.763 0.816 +0.053
Aspect 0.717 0.803 +0.086
Aspect Sentiment 0.757 0.600 0.157

Three tasks improved substantially. Aspect Sentiment collapsed. Bug-report keyword reviews were heavily negative skewed, so oversampling them shifted the training distribution to 80.9% negative sentiment, in an MTL scenario this negatively affects the ability to learn the shared representations.

In a shared-encoder MTL setup, data curation for one task is never just curation for that task alone. Shared Learning is also shared interference.

(The models were not cross-evaluated - see Limitations.)

Dataset

Manually annotated from the Ola Vs Uber Reviews dataset using a custom Tkinter annotation tool (multitag.py) with keyboard shortcuts and enforced label completeness.

Stage Count
Raw corpus 1,069,616
After cleaning + ≥5-word filter 495,036
Original distribution, annotated 4,999
Keyword-boosted, annotated 4,997
Split 70% train / 15% val / 15% test

Class imbalance is severe in the wild: bug reports are 18.4% of the original set, neutral sentiment just 5.2%. Handled via inverse-frequency class weighting during training, not by rebalancing the raw data.

Data availability: the annotated datasets (original and keyword-boosted, train/val/test splits) are included in this repository under data/. Raw source: Uber Customer Reviews on Kaggle.

Repository structure

.
├── architecture.png    # Overview diagram
├── environment.yml     # conda environment spec
├── data/               # annotated original + keyword-boosted splits
├── notebooks/          # preprocessing, annotation QA, analysis, inference prep
├── src/
│   ├── preprocess.py   # cleaning + filtering pipeline
│   ├── sampler.py      # original / keyword-boosted sampling strategies
│   ├── multitag.py     # Tkinter annotation tool
│   ├── dataset.py      # tokenisation + PyTorch Dataset
│   ├── model.py        # shared-encoder MTL + single-task model definitions
│   ├── train.py        # training loop (MTL + STL)
│   ├── evaluate.py     # Macro-F1, confusion matrices, McNemar's test
│   └── infer.py        # inference on new review text
└── README.md

Running the pipeline

0. Optional Setup and get data

cd <project location> # replace with location
mkdir -p data/
mkdir -p data/raw/
mkdir -p data/processed/
mkdir -p outputs/
mkdir -p runs/ # if using tensorboard
mkdir -p outputs/inference/

# Optionally replace with own data though could cause slight issues
curl -L -o data/raw/ola-vs-uber-play-store-reviews.zip\<newline>
  https://www.kaggle.com/api/v1/datasets/download/khushipitroda/ola-vs-uber-play-store-reviews

cd data/raw
unzip ola-vs-uber-play-store-reviews.zip
mv "Uber Customer Reviews.csv" uber_reviews.csv
rm ola-vs-uber-play-store-reviews.zip
rm “Ola Customer Reviews.csv”
cd ../..

conda env create -f environment.yml
conda activate reclass

1. Preprocess raw reviews (clean, filter, sample)

python src/preprocess.py [optionally fill in INPUT and OUTPUT at top of file]

2. Sample cleaned reviews for tagging

python src/sampler.py [optionally fill in target_samples, INPUT and OUTPUT at top of file]

3. Annotate (or use existing labelled CSVs)

python src/multitag.py [optionally fill in INPUT and OUTPUT at top of file]

3.5 Train Test Split

Use the notebook preprocessing_tagged.ipynb to create splits, cells 2 and 23 contain input/output paths.

  • Inputs data/raw/uber_reviews_tagged_original.csv and /data/raw/uber_reviews_tagged_boosted.csv by default.
  • Outputs to data/tagged_boosted_cleaned.csv and data/tagged_original_cleaned.csv by default.

4. Train - MTL or single-task baseline

python src/train.py 
Defaults to python src/train.py --mode mtl --task all --dataset original batch_size 16 --epochs 5 --patience 3 --lr 2e-5

5. Evaluate against test set

python src/evaluate.py

options:
  -h, --help            show this help message and exit
  --mode {mtl,stl}      mtl or stl
  --task {all,bug_report,feature_request,aspect,aspect_sentiment}
  --dataset {original,boosted}
  --model_path MODEL_PATH .pt file path
  --batch_size BATCH_SIZE

e.g. python src/evaluate.py --mode mtl --dataset original --model_path outputs/best_model_mtl_original.pt

6. Run inference on new text

python src/infer.py
options:
  -h, --help            show this help message and exit
  --model MODEL         .pt file - just the filename not path
  --task {all,bug_report,feature_request,aspect,aspect_sentiment}
  --interactive         Loops reading input until exit()
  --text                Use command line text for input
  --dataset DATASET     Enter a file name for inference (stored in data/processed/)
  --batch_size BATCH_SIZE
  --mode {mtl,stl}      		mtl or stl
  --text_column TEXT_COLUMN 		Where is the text column

e.g. python src/infer.py --model best_model_mtl_original.pt --mode mtl --text  

Trained on a single NVIDIA RTX 2070 Super (8GB VRAM), ~12 GPU-hours total across all reported runs.

Limitations

  • Single annotator. No inter-annotator agreement / Cohen's Kappa, so inconsistencies / bias can't be ruled out.
  • No cross-evaluation. The boosted model was only tested on its own (curated - not original distribution) test set, so its performance on the real data was left unknown.
  • Single domain. Cross domain performance was left unknown.
  • Overconfidence on errors. Overconfidence was too high and needs calibrating in the future.
  • Linear classification heads. Tasks with multiple classes would likely benefit from an additional layer.

Future work

  • Cross-evaluate the boosted model on the original test set, quickest next step.
  • Second annotator on any future data labelling.
  • Dynamic or learned weighting.
  • Possibly implement the transferrable elements elsewhere on future projects.
  • Temperature scaling on the overconfidence.

Citation / provenance

Built as a BSc Computer Science dissertation project, Oxford Brookes University (2026). Uses the Ola Vs Uber Reviews also uploaded here Ola Vs Uber Play store reviews from Kaggle, and XLM-RoBERTa (Conneau et al., 2020) via HuggingFace Transformers.

License

MIT - see LICENSE

Description
Review Classification using XLM-RoBERTa in a multitask configuration to compare against single head baselines. All Data, statistics and model checkpoints used are attached.
Readme MIT 6.7 GiB
Languages
Jupyter Notebook 68.7%
Python 30.1%
Batchfile 1.2%