This repository contains the official implementation of the following paper:
Pyo, S., Kim, S., Li, X., & Kim, J. (2026). Beyond Sentiment: A Dual-Channel Approach to Emotion and Review Summarization for Recommendation. IEEE Access.
DCESR (Dual-Channel Emotion- & Semantic-aware Recommender) is an advanced recommendation framework designed to bridge the gap between textual feedback and rating prediction. Unlike traditional models, DCESR operates through two distinct analytical channels to capture the full spectrum of user-item interactions. The Semantic Channel utilizes the BART model for abstractive summarization, extracting high-level item attributes and long-term user preferences from noisy review text. Simultaneously, the Emotion Channel employs DistilRoBERTa to distill 7-class emotional probability vectors, capturing the psychological nuances behind user ratings. These heterogeneous features are then adaptively integrated through a Gated Multimodal Unit (GMU), which learns a dynamic weighting mechanism to balance the importance of semantic context and emotional sentiment for each specific prediction. By fusing deep linguistic understanding with fine-grained sentiment analysis, DCESR provides a robust and interpretative solution for modern recommender systems.
This project is implemented in Python 3.8+. To ensure reproducibility, please install the specific versions of the libraries listed below.
| Category | Library | Version | Description |
|---|---|---|---|
| Deep Learning | TensorFlow / Keras |
2.21.0 / 3.13.2 |
Implements the core neural network and Gated Multimodal Unit (GMU). |
| NLP | Transformers |
5.3.0 |
Provides pre-trained BART (Semantics) and DistilRoBERTa (Emotion) models. |
| NLP Backend | PyTorch |
2.11.0 |
Serves as the high-performance backend for Hugging Face transformer models. |
| Analysis | Pandas |
3.0.1 |
Handles data loading, 5-core filtering, and review set aggregation. |
| Matrix | NumPy |
2.4.3 |
Facilitates efficient numerical operations on 768D and 7D embedding vectors. |
| ML Tools | scikit-learn |
1.8.0 |
Manages train/val/test splitting and calculates performance metrics (MAE, RMSE). |
PyYAML(6.0.3): Essential for parsing theconfig.yamlto manage hyperparameters and file paths dynamically.PyArrow(23.0.1): Used as the high-performance engine for saving and loading large Parquet data splits.tqdm(4.67.3): Provides real-time visual feedback for long-running embedding extraction processes.h5py(3.14.0): Handles the serialization and storage of trained Keras model weights.Huggingface Hub(1.7.2): Manages the seamless downloading of pre-trained NLP model weights.
The repository is organized as follows to ensure a clear workflow from data preprocessing to model evaluation:
DCESR/
├── main.py # Central orchestrator to run the entire pipeline
├── config.yaml # Global configurations (hyperparameters, paths, device)
├── requirements.txt # List of required Python libraries
├── .gitignore # Specifies files/folders to be ignored by Git
├── README.md # Project documentation and overview
│
├── model/ # Model Architecture
│ ├── __init__.py
│ └── proposed.py # DCESR Model (BART + DistilRoBERTa + GMU)
│
├── scr/ # Source Code Modules
│ ├── __init__.py
│ ├── data_processing.py # Data loading, 5-core filtering, and set aggregation
│ ├── trainer.py # Training loop, evaluation metrics, and data splitting
│ ├── bart.py # BART-based semantic feature extraction
│ └── distilroberta.py # DistilRoBERTa-based emotion feature extraction
│
└── data/ # Data Directory
├── raw/
│ └── SampleData.json.gz # Sample dataset for testing
└── processed/ # (Generated) Processed .pkl and .parquet files
We recommend using a virtual environment to manage dependencies.
python -m venv .venv
pip install -r requirements.txtThe model requires the dataset to be placed in the specific directory defined in the structure.
- Prepare Data: Place your original dataset (e.g., SampleData.json.gz) into the data/raw/ directory.
- Automatic Preprocessing: When you run the main script, the pipeline will automaticall performs.
You can customize hyperparameters and file paths in the centralized config file.
- File Path:
config.yaml
Once the environment and data are ready, execute the following command to start the full workflow (Preprocessing → Training → Evaluation):
python main.py
The DCESR (Dual-Channel Emotion- & Semantic-aware Recommender) is a deep learning-based recommendation framework designed to predict user ratings by analyzing the multidimensional nature of review texts.
The model operates through a dual-channel architecture. The Semantic Channel utilizes a pre-trained BART encoder to extract 768-dimensional vectors from summarized user and item review sets, capturing high-level attributes and preferences. Simultaneously, the Emotion Channel employs DistilRoBERTa to analyze 7-class emotion probabilities, generating a fine-grained psychological profile for each user and item.
To integrate these diverse features, the model uses a Gated Multimodal Unit (GMU), which adaptively calculates weights to balance the importance of semantic meaning versus emotional sentiment for every specific prediction. Finally, this fused representation is passed through a neural network to output a precise numerical rating.
The following table summarizes the performance comparison across three Amazon review datasets. DCESR consistently achieves the lowest error rates in both MAE and RMSE.
| Model | Books | Movies and TV | Office Products | |||
|---|---|---|---|---|---|---|
| MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | MAE ↓ | RMSE ↓ | |
| NCF | 0.703 | 0.902 | 1.099 | 1.360 | 0.638 | 0.911 |
| DeepCoNN | 0.525 | 0.767 | 0.784 | 1.086 | 0.580 | 0.855 |
| NARRE | 0.530 | 0.791 | 0.780 | 1.083 | 0.584 | 0.858 |
| DRRNN | 0.498 | 0.764 | 0.724 | 1.081 | 0.523 | 0.844 |
| AENAR | 0.508 | 0.760 | 0.750 | 1.083 | 0.559 | 0.852 |
| MFNR | 0.501 | 0.777 | 0.738 | 1.093 | 0.526 | 0.850 |
| DCESR (Ours) | 0.485 | 0.745 | 0.690 | 1.035 | 0.498 | 0.809 |